Papers with Large Audio Language Models
Audio Query Handling System with Integrated Expert Models and Contextual Understanding (2025.emnlp-industry)
Copied to clipboard
| Challenge: | Existing chatbots are limited to specific audio tasks, but the domain of audio content related queries remains underexplored. |
| Approach: | They propose to use an intent classifier to route queries to audio-related experts using a diverse audio query dataset. |
| Outcome: | The proposed system outperforms state-of-the-art LLMs on custom audio tasks and MMAU sound set benchmarks. |
Enhancing Temporal Understanding in Audio Question Answering for Large Audio Language Models (2025.naacl-industry)
Copied to clipboard
| Challenge: | Recent literature focuses on constructing large audio language models (LALMs) but they are limited in temporal reasoning, which may hinder commercial applications . |
| Approach: | They propose a data augmentation technique for generating reliable audio temporal questions and answers using an LLM. |
| Outcome: | The proposed model performs well on public audio benchmark datasets and is optimized for edge applications. |
SCENEBench: An Audio Understanding Benchmark Grounded in Assistive and Industrial Use Cases (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing models that measure audio comprehension beyond automatic speech recognition lack performance and latency. |
| Approach: | They propose a benchmark suite that measures audio comprehension beyond automatic speech recognition . the benchmark suite includes a small human-recorded evaluation split per category . |
| Outcome: | The proposed suite measures audio comprehension beyond speech recognition . it includes a small human-recorded evaluation split per category . |
Reshaping Representation Space to Balance the Safety and Over-rejection in Large Audio Language Models (2025.emnlp-main)
Copied to clipboard
| Challenge: | Large Audio Language Models (LALMs) have demonstrated unprecedented capabilities in natural language understanding and generation, revolutionizing human-machine dialogue. |
| Approach: | They propose an unsupervised safety-fine-tuning strategy that reshapes LALMs representation space to enhance existing LALM safety-alignment while balancing the risk of over-rejection. |
| Outcome: | The proposed approach improves LALMs safety under three input conditions while increasing over-rejection rate by only 0.88% on average. |
SEE: Signal Embedding Energy for Quantifying Noise Interference in Large Audio Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing studies on noise lack quantitative analysis and rely on intuition and empirical observation, thus failing to understand practical robustness. |
| Approach: | They propose a method for quantifying the impact of noise intensity on LALM inputs by using a structured activation subspace derived from the model's internal representations. |
| Outcome: | The proposed method outperforms existing denoising methods and demonstrates that noise is perceived more accurately than raw audio features. |
CORD: Bridging the Audio–Text Reasoning Gap via Weighted On-policy Cross-modal Distillation (2026.findings-acl)
Copied to clipboard
Hu Jing, Danxiang Zhu, Xianlong Luo, Dan Zhang, Shuwei He, Yishu Lei, Shikun Feng, Hai-Tao Zheng, Jingzhou HE, Yu Sun, Hua Wu, Haifeng Wang
| Challenge: | Large Audio Language Models (LALMs) exhibit a degradation in knowledge and reasoning capabilities . empirical results show that CORD significantly bridges the audio–text performance gap . |
| Approach: | They propose a framework that performs online cross-modal self-distillation to bridge the acoustic-semantic gap between LALMs and text-based models. |
| Outcome: | The proposed framework bridges the acoustic-semantic gap between LALMs and text-based models . it employs on-policy reverse KL divergence with importance-aware weighting . |
Think Smart, Not Hard: Difficulty Adaptive Reasoning for Large Audio Language Models (2026.findings-acl)
Copied to clipboard
| Challenge: | Existing methods to determine whether to perform reasoning lack fine-grained mechanisms to adapt reasoning length to problem complexity. |
| Approach: | They propose a difficulty-adaptive reasoning method that dynamically links reasoning length to the model’s perceived problem difficulty. |
| Outcome: | The proposed method reduces average reasoning length by 50%, achieving higher efficiency without sacrificing accuracy. |
Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding (2026.findings-acl)
Copied to clipboard
| Challenge: | Recent Large Audio Language Models (LALMs) have shown strong capabilities in audio understanding, yet their reasoning remains vulnerable to perceptual errors. |
| Approach: | They propose a large-scale dataset for **Perception-Aware Question Answering** that uses a hierarchical decoupling strategy to separate speech from environmental sounds and distinguishes among multiple speakers. |
| Outcome: | The proposed model improves on MMAU-mini, MMAR, and PAQA while maintaining comparable performance on multiple benchmarks. |
PolyAudio: Advancing Multi-Audio Reasoning in Large Audio Language Models with Interleaved Multi-Audio Contexts (2026.findings-acl)
Copied to clipboard
Sonal Kumar, Sreyan Ghosh, Yueqian Lin, S Sakshi, Ashish Seth, Yiran Chen, Ramani Duraiswami, Dinesh Manocha
| Challenge: | Large Audio Language Models have shown impressive performance on single-clip tasks . however, their ability to reason over interleaved multi-audio contexts remains limited . |
| Approach: | They propose a LALM that targets multi-audio understanding via instruction tuning rather than massive-scale pre-training. |
| Outcome: | The proposed model outperforms baseline models on multi-audio tasks while maintaining robustness. |